NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge

https://doi.org/https://doi.org/10.48550/arXiv.2407.05941

Eliopoulos, Nick John; Jajal, Purvish; Davis, James; Liu, Gaowen; Thiravathukal, George; Lu, YungHsiang (February 2025, The Computer Vision Foundation.)

This paper investigates how to efficiently deploy vision transformers on edge devices for small workloads. Recent methods reduce the latency of transformer neural networks by removing or merging tokens, with small accuracy degradation. However, these methods are not designed with edge device deployment in mind: they do not leverage information about the latency-workload trends to improve efficiency. We address this shortcoming in our work. First, we identify factors that affect ViT latency-workload relationships. Second, we determine token pruning schedule by leveraging non-linear latency-workload relationships. Third, we demonstrate a training-free, token pruning method utilizing this schedule. We show other methods may increase latency by 2-30%, while we reduce latency by 9-26%. For similar latency (within 5.2% or 7ms) across devices we achieve 78.6%-84.5% ImageNet1K accuracy, while the state-of-the-art, Token Merging, achieves 45.8%-85.4%.
more » « less
Free, publicly-accessible full text available February 28, 2026
Detecting Music Performance Errors with Transformers

https://doi.org/10.1609/aaai.v39i22.34539

Chou, Benjamin Shiue_Hal; Jajal, Purvish; Eliopoulos, Nicholas John; Nadolsky, Tim; Yang, Cheng_Yun; Ravi, Nikita; Davis, James C; Yun, Kristen Yeon_Ji; Lu, Yung_Hsiang (April 2025, Proceedings of the AAAI Conference on Artificial Intelligence)

Beginner musicians often struggle to identify specific errors in their performances, such as playing incorrect notes or rhythms. There are two limitations in existing tools for music error detection: (1) Existing approaches rely on automatic alignment; therefore, they are prone to errors caused by small deviations between alignment targets; (2) There is insufficient data to train music error detection models, resulting in over-reliance on heuristics. To address (1), we propose a novel transformer model, Polytune, that takes audio inputs and outputs annotated music scores. This model can be trained end-to-end to implicitly align and compare performance audio with music scores through latent space representations. To address (2), we present a novel data generation technique capable of creating large-scale synthetic music error datasets. Our approach achieves a 64.1% average Error Detection F1 score, improving upon prior work by 40 percentage points across 14 instruments. Additionally, our model can handle multiple instruments compared with existing transcription methods repurposed for music error detection.
more » « less
Free, publicly-accessible full text available April 11, 2026
Token Turing Machines are Efficient Vision Models

https://doi.org/10.48550/arXiv.2409.07613

Jajal, Purvish; Eliopoulos, Nick John; Chou, Benjamin Shiue-Hal; Thiruvathukal, George K; Davis, James C; Lu, Yung-Hsiang (February 2025, The Computer Vision Foundation.)

We propose Vision Token Turing Machines (ViTTM), an efficient, low-latency, memory-augmented Vision Transformer (ViT). Our approach builds on Neural Turing Machines and Token Turing Machines, which were applied to NLP and sequential visual understanding tasks. ViTTMs are designed for non-sequential computer vision tasks such as image classification and segmentation. Our model creates two sets of tokens: process tokens and memory tokens; process tokens pass through encoder blocks and read-write from memory tokens at each encoder block in the network, allowing them to store and retrieve information from memory. By ensuring that there are fewer process tokens than memory tokens, we are able to reduce the inference time of the network while maintaining its accuracy. On ImageNet-1K, the state-of-the-art ViT-B has median latency of 529.5ms and 81.0% accuracy, while our ViTTM-B is 56% faster (234.1ms), with 2.4 times fewer FLOPs, with an accuracy of 82.9%. On ADE20K semantic segmentation, ViT-B achieves 45.65mIoU at 13.8 frame-per-second (FPS) whereas our ViTTM-B model acheives a 45.17 mIoU with 26.8 FPS (+94%).
more » « less
Free, publicly-accessible full text available February 28, 2026
Pruning One More Token is Enough: Leveraging Latency-Workload Non-Linearities for Vision Transformers on the Edge

https://doi.org/10.1109/WACV61041.2025.00695

Eliopoulos, Nicholas John; Jajal, Purvish; Davis, James C; Liu, Gaowen; Thiravathukal, George K; Lu, Yung-Hsiang (February 2025, IEEE)

Free, publicly-accessible full text available February 26, 2026
Token Turing Machines are Efficient Vision Models

https://doi.org/10.1109/WACV61041.2025.00767

Jajal, Purvish; Eliopoulos, Nick John; Chou, Benjamin Shiue-Hal; Thiravathukal, George K; Davis, James C; Lu, Yung-Hsiang (February 2025, IEEE)

Free, publicly-accessible full text available February 26, 2026
Securing Deep Neural Networks on Edge from Membership Inference Attacks Using Trusted Execution Environments

Yang, Cheng-Yun; Ramshankar, Gowri; Nambiar, Sudarshan; Miller, Evan; Zhang, Xun; Eliopoulos, Nicholas; Jajal, Purvish; Jing_Tian, Dave; Chen, Shuo-Han; Perng, Chiy-Ferng; et al (August 2024, 2024 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED))

Full Text Available
An automated approach for improving the inference latency and energy efficiency of pretrained CNNs by removing irrelevant pixels with focused convolutions

https://doi.org/10.1109/ASP-DAC58780.2024.10473884

Tung, Caleb; Eliopoulos, Nicholas; Jajal, Purvish; Ramshankar, Gowri; Yang, Cheng-Yun; Synovic, Nicholas; Zhang, Xuecen; Chaudhary, Vipin; Thiruvathukal, George K; Lu, Yung-Hsiang (January 2024, Asia and South Pacific Design Automation Conference (ASP-DAC))

Computer vision often uses highly accurate Convolutional Neural Networks (CNNs), but these deep learning models are associated with ever-increasing energy and computation requirements. Producing more energy-efficient CNNs often requires model training which can be cost-prohibitive. We propose a novel, automated method to make a pretrained CNN more energyefficient without re-training. Given a pretrained CNN, we insert a threshold layer that filters activations from the preceding layers to identify regions of the image that are irrelevant, i.e. can be ignored by the following layers while maintaining accuracy. Our modified focused convolution operation saves inference latency (by up to 25%) and energy costs (by up to 22%) on various popular pretrained CNNs, with little to no loss in accuracy
more » « less
Full Text Available
Reusing Deep Learning Models: Challenges and Directions in Software Engineering

https://doi.org/10.1109/JVA60410.2023.00015

Davis, James C; Jajal, Purvish; Jiang, Wenxin; Schorlemmer, Taylor R; Synovic, Nicholas; Thiruvathukal, George K (July 2023, IEEE)

Full Text Available
PTMTorrent: A Dataset for Mining Open-source Pre-trained Model Packages

https://doi.org/10.1109/MSR59073.2023.00021

Jiang, Wenxin; Synovic, Nicholas; Jajal, Purvish; Schorlemmer, Taylor R; Tewari, Arav; Pareek, Bhavesh; Thiruvathukal, George K; Davis, James C (May 2023, IEEE)
Interoperability in Deep Learning: A User Survey and Failure Analysis of ONNX Model Converters

https://doi.org/10.1145/3650212.3680374

Jajal, Purvish; Jiang, Wenxin; Tewari, Arav; Kocinare, Erik; Woo, Joseph; Sarraf, Anusha; Lu, Yung-Hsiang; Thiruvathukal, George K; Davis, James C (September 2024, ACM -- ISSTA)

Full Text Available

Search for: All records